Papers with pointwise training
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling (2026.acl-long)
Copied to clipboard
| Challenge: | Existing reward models lack generative and reasoning capabilities, resulting in poor performance. |
| Approach: | They propose a reward-aware task-adaptive reward model that enables pointwise training using readily available pairwise data via a novel Preference-Aware Reward mechanism. |
| Outcome: | The proposed reward model achieves an average relative improvement of 8.7% over the base models on RewardBench and RMBench. |
Mitigating Selection Bias in Large Language Models via Permutation-Aware GRPO (2026.acl-long)
Copied to clipboard
| Challenge: | Existing inference-time debiasing ignores that the same question should yield consistent answers across permutations. |
| Approach: | They propose a permutation-aware group-relative policy optimization which enforces permutations-consistent semantic reasoning. |
| Outcome: | The proposed model outperforms strong baselines across seven benchmarks while maintaining high overall performance. |